Predicting agent trajectories

KR1020260122949APending Publication Date: 2026-08-12MOTIONAL AD LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2026-08-12

Smart Images

  • Figure P1020267024910_ABST
    Figure P1020267024910_ABST
Patent Text Reader

Abstract

A method for predicting agent trajectories is provided, the method may include the steps of generating a graph corresponding to a map of a scene by encoding map features and agent features as node encodings of the graph, and determining a policy to apply to edges exiting from nodes of the graph. Some of the described methods also include the steps of sampling a path for a target vehicle in a scene according to the policy, and predicting a set of trajectories based on the sampled path traversed by the policy and the sampled latent variables. A system and a computer program product are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference regarding related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 179,169, filed on April 23, 2021, titled “Predicting Agent Trajectories”. The disclosure of the aforementioned application is incorporated herein by reference in its entirety for all purposes. Background Technology

[0003] There is inherent uncertainty in predicting the future, which makes trajectory prediction a difficult problem. To navigate complex traffic scenes safely and efficiently, it is desirable for autonomous vehicles to accurately predict the future trajectories of surrounding vehicles and thus plan their own trajectories accordingly. Brief explanation of the drawing

[0004] FIG. 1 is an exemplary environment in which a vehicle including one or more components of an autonomous driving system can be implemented. Figure 2 is a diagram of one or more systems of a vehicle including an autonomous driving system. FIG. 3 is a diagram of the components of one or more devices and / or one or more systems of FIG. 1 and FIG. 2. Figure 4a is a diagram of specific components of an autonomous driving system. Figure 4b is a diagram of the implementation of a neural network. Figures 4c and 4d are diagrams illustrating the exemplary operation of a CNN. Figure 5 is a block diagram illustrating an exemplary process of agent trajectory prediction. Figure 6 is a block diagram illustrating an exemplary agent trajectory prediction system. Figure 7 illustrates an exemplary graph encoder. FIG. 8 illustrates an exemplary HD map and an exemplary directed graph corresponding to the exemplary HD map. Figure 9 illustrates an exemplary policy header. Figure 10 illustrates an exemplary trajectory decoder. FIGS. 11a through 11c each illustrate three exemplary directed graphs generated by a lane graph generator. FIG. 12 illustrates the aggregation of agent context and map context according to an implementation of the present disclosure. Figure 13 illustrates an exemplary reward model that can be used in a policy header. Figure 14 illustrates an exemplary process for predicting a trajectory for an agent. Specific details for implementing the invention

[0005] In the following description, numerous specific details are described for illustrative purposes to provide a complete understanding of the present disclosure. However, it will be apparent that the embodiments described by the present disclosure can be practiced without these specific details. In some cases, well-known structures and devices are illustrated in block diagram form to avoid unnecessarily obscuring the aspects of the present disclosure.

[0006] Specific arrangements or sequences of schematic elements, such as those representing systems, devices, modules, instruction blocks, data elements, etc., are illustrated in the drawings for convenience of explanation. However, a person skilled in the art will understand that a specific order or arrangement of schematic elements in the drawings does not imply that a specific processing order or sequence of processes, or a separation of processes, is required unless explicitly stated otherwise. Furthermore, the inclusion of schematic elements in the drawings does not imply that, in some embodiments, such elements are required in all embodiments, or that features represented by such elements may not be included in other elements or combined with other elements, unless explicitly stated otherwise.

[0007] Furthermore, where connecting elements, such as solid or dashed lines or arrows, are used in the drawings to illustrate a connection, relationship, or association between two or more different schematic elements, the absence of any such connecting elements does not imply that a connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements are not illustrated in the drawings to avoid obscuring the present disclosure. Additionally, for convenience of illustration, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, where a connecting element represents the communication of signals, data, or instructions (e.g., "software instructions"), a person skilled in the art will understand that such an element may represent one or more signal paths (e.g., buses) that may be required to perform the communication.

[0008] Terms such as first, second, third, etc. are used to describe various components, but these elements should not be limited by these terms. Terms such as first, second, third, etc. are used solely to distinguish one element from another. For example, without departing from the scope of the described embodiments, the first contact may be referred to as the second contact, and similarly, the second contact may be referred to as the first contact. Both the first contact and the second contact are contacts, but they are not the same contact.

[0009] The technical terms used in the description of the various described embodiments herein are included solely to describe specific embodiments and are not intended to be limiting. As used in the description of the various described embodiments and in the appended claims, singular forms (“a,” “an,” and “the”) are intended to include plural forms and may be used interchangeably with “one or more” or “at least one” unless the context otherwise clearly indicates. It will also be understood that the term “and / or,” as used herein, refers to and includes parts of one or more of the associated enumerated items and all possible combinations thereof. Furthermore, it will be understood that the terms “include” and / or “include,” when used herein, specify the presence of the mentioned features, integers, steps, actions, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof.

[0010] As used herein, the terms “communication” and “to communicate” refer to at least one of the reception, receipt, transmission, delivery, provision, etc. of information (or, for example, information expressed by data, signals, messages, instructions, commands, etc.). One unit (for example, a device, a system, a component of a device or system, a combination thereof, etc.) communicating with another unit means that one unit can receive information directly or indirectly from another unit and / or transmit information to the other unit (for example, transmit). This may refer to a direct or indirect connection that is essentially wired and / or wireless. Additionally, two units may be communicating with each other even if the transmitted information may be modified, processed, relayed, and / or routed between the first unit and the second unit. For example, the first unit may be communicating with the second unit even if the first unit passively receives information and does not actively transmit information to the second unit. As another example, if at least one intermediate unit (e.g., a third unit located between the first unit and the second unit) processes information received from the first unit and transmits the processed information to the second unit, the first unit may be communicating with the second unit. In some embodiments, the message may refer to a network packet containing data (e.g., a data packet, etc.).

[0011] As used herein, the term “in case” is optionally interpreted, depending on the context, to mean “when” or “at” or “in response to a decision” or “in response to detecting”. Similarly, the phrase “when it is decided” or “when [the mentioned condition or event] is detected” is optionally interpreted, depending on the context, to mean “when it is decided”, “in response to a decision”, “when [the mentioned condition or event] is detected”, or “in response to detecting [the mentioned condition or event]”. Furthermore, as used herein, terms such as “has, have” and “having” are intended to be open-ended terms. In addition, the phrase “based on” is intended to mean “at least partially based on” unless otherwise explicitly stated.

[0012] Examples of embodiments illustrated in the attached drawings will now be described in detail. In the following detailed description, many specific details are provided to provide a complete understanding of the various described embodiments. However, it will be apparent to those skilled in the art that the various described embodiments can be practiced without these specific details. In other cases, known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure the aspects of the embodiments.

[0013] General Overview

[0014] In some aspects and / or embodiments, the system, method, and computer program products described herein include and / or implement using a graphical representation of a scene to predict agent trajectories. A vehicle (e.g., an autonomous vehicle) navigates through the traffic scene (e.g., within the environment) by predicting the future trajectories of surrounding agents (e.g., target vehicles or pedestrians). High-resolution maps and agent trajectories are encoded into a graph representing the scene. Agent trajectories are predicted according to discrete policies (e.g., state-action mapping), and for each trajectory, context is optionally aggregated (e.g., the context is conditioned on lane graph traversal). Latent variables are sampled to account for the longitudinal variability of the predicted trajectories. The final set of trajectories for each agent in the scene is determined by sampling the path (root) traversed by the policy and the latent distribution (e.g., K-means clustering, output cluster center).

[0015] Some of the advantages of these techniques include selective conditional prediction for lane graph traversal, which generates trajectories that are (i) diverse in terms of routes to create most possible paths based on the data, and (ii) precise and scene-compliant with the lowest off-road rate. Precise trajectories refer to paths that are likely to occur based on the data. Using sampled latent variables generates diverse trajectories in terms of motion profiles. Furthermore, this technique improves sampling efficiency through graph-based representations with individual trajectories. Context is specifically captured at the nodes, ensuring the inclusion of relevant context for each trajectory. In some cases, existing techniques fail to capture relevant context by using convolutional neural network receptive fields. Additionally, this technique is inevitably rooted in the computing technology that enables the operation of autonomous vehicles. The agent model in the scene is improved, resulting in more accurate prediction of agent behavior. Since the graph location consists of drivable areas where undrivable regions appear infrequently, the model is efficient in terms of state space. This allows for more efficient use of computing resources, including processing and memory resources.

[0016] Now, referring to FIG. 1, an exemplary environment (100) is illustrated in which vehicles including autonomous driving systems as well as vehicles not including autonomous driving systems are operated. As illustrated, the environment (100) includes vehicles (102a to 102n), objects (104a to 104n), routes (106a to 106n), a region (108), a vehicle-to-infrastructure (V2I) device (110), a network (112), a remote autonomous vehicle (AV) system (114), a fleet management system (116), and a V2I system (118). Vehicles (102a to 102n), vehicle-to-infrastructure (V2I) devices (110), networks (112), autonomous vehicle (AV) systems (114), fleet management systems (116), and V2I systems (118) are interconnected via wired connections, wireless connections, or a combination of wired and wireless connections (e.g., establishing connections for communication, etc.). In some embodiments, objects (104a to 104n) are interconnected with at least one of the vehicles (102a to 102n), vehicle-to-infrastructure (V2I) devices (110), networks (112), autonomous vehicle (AV) systems (114), fleet management systems (116), and V2I systems (118) via wired connections, wireless connections, or a combination of wired and wireless connections.

[0017] Vehicles (102a through 102n) (individually referred to as vehicles (102) and collectively referred to as vehicles (102)) comprise at least one device configured to transport goods and / or people. In some embodiments, vehicles (102) are configured to communicate with a V2I device (110), a remote AV system (114), a fleet management system (116), and / or a V2I system (118) via a network (112). In some embodiments, vehicles (102) include automobiles, buses, trucks, trains, etc. In some embodiments, vehicles (102) are identical or similar to the vehicles (200) described herein (see FIG. 2). In some embodiments, a vehicle (200) among a set of vehicles (200) is associated with an autonomous fleet manager. In some embodiments, the vehicles (102) travel along their respective routes (106a to 106n) as described herein (individually referred to as route (106) and collectively referred to as routes (106)). In some embodiments, one or more vehicles (102) include an autonomous driving system (e.g., an autonomous driving system identical or similar to the autonomous driving system (202)).

[0018] Objects (104a to 104n) (individually referred to as object (104) and collectively referred to as object (104)) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, at least one structure (e.g., a building, a sign, a fire hydrant, etc.). Each object (104) is stationary (e.g., located in a fixed position for a period of time) or moving (e.g., having speed and associated with at least one trajectory). In some embodiments, the objects (104) are associated with corresponding locations within the area (108).

[0019] Routes (106a through 106n) (individually referred to as Route (106) and collectively referred to as Route (106)) are each associated with a sequence of actions (also referred to as a trajectory) connecting states in which the AV can operate (e.g., defines this). Each route (106) starts in an initial state (e.g., a state corresponding to a first spatiotemporal position, velocity, etc.) and a final target state (e.g., a state corresponding to a second spatiotemporal position different from the first spatiotemporal position) or a target area (e.g., a subspace of acceptable states (e.g., terminal states). In some embodiments, the first state includes a location where individuals or individuals must be picked up by the AV, and the second state or area includes a location or locations where individuals or individuals picked up by the AV must be dropped off. In some embodiments, the routes (106) include a plurality of acceptable state sequences (e.g., a plurality of spatiotemporal position sequences), and the plurality of state sequences are associated with a plurality of trajectories (e.g., defines them). In one example, the routes (106) include only higher-level actions or inaccurate state positions, such as a series of connected roads indicating turning directions at road intersections. Additionally or alternatively, the routes (106) may include more accurate actions or states, such as, for example, accurate positions within specific target lanes or lane regions and target speeds at said positions.In one example, the routes (106) include multiple exact state sequences following at least one upper-level action sequence having a limited lookahead horizon to reach intermediate goals, wherein a combination of consecutive repetitions of the limited horizon state sequences is accumulated to correspond to multiple trajectories, and these multiple trajectories collectively form an upper-level route that terminates in a final goal state or region.

[0020] Area (108) includes a physical area (e.g., a geographical area) over which vehicles (102) can travel. In one example, the area (108) includes at least one state (e.g., a country, a province, an individual state of a plurality of states included in a country), at least one part of a state, at least one city, at least one part of a city, etc. In some embodiments, the area (108) includes at least one named thoroughfare (referred to herein as “road”), such as a main road, an interstate main road, a park road, a city street, etc. Additionally or alternatively, in some examples, the area (108) includes at least one unnamed road, such as an access road, a section of a parking lot, a section of an open space and / or undeveloped site, an unpaved path, etc. In some embodiments, the road includes at least one lane (e.g., a part of the road that can be traversed by a vehicle (102)). In one example, the road includes at least one lane associated with at least one lane marking (e.g., identified based thereon).

[0021] A vehicle-to-infrastructure (V2I) device (110) (sometimes referred to as a vehicle-to-infrastructure (V2X) device) comprises at least one device configured to communicate with vehicles (102) and / or a V2I infrastructure system (118). In some embodiments, the V2I device (110) is configured to communicate with vehicles (102), a remote AV system (114), a fleet management system (116), and / or a V2I system (118) via a network (112). In some embodiments, the V2I device (110) includes a radio frequency identification (RFID) device, signage, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), a lane marker, a street light, a parking meter, etc. In some embodiments, the V2I device (110) is configured to communicate directly with vehicles (102). Additionally or alternatively, in some embodiments, the V2I device (110) is configured to communicate with vehicles (102), a remote AV system (114), and / or a fleet management system (116) via a V2I system (118). In some embodiments, the V2I device (110) is configured to communicate with the V2I system (118) via a network (112).

[0022] The network (112) includes one or more wired and / or wireless networks. In one example, the network (112) includes a cellular network (e.g., LTE (long term evolution) network, 3G (third generation) network, 4G (fourth generation) network, 5G (fifth generation) network, CDMA (code division multiple access) network, etc.), PLMN (public land mobile network), LAN (local area network), WAN (wide area network), MAN (metropolitan area network), telephone network (e.g., PSTN (public switched telephone network)), private network, ad-hoc network, intranet, internet, fiber-optic-based network, cloud computing network, etc., and a combination of some or all of these networks.

[0023] The remote AV system (114) includes at least one device configured to communicate with vehicles (102), a V2I device (110), the network (112), the remote AV system (114), a fleet management system (116), and / or a V2I system (118) via a network (112). In one example, the remote AV system (114) includes a server, a group of servers, and / or other similar devices. In some embodiments, the remote AV system (114) is co-located with the fleet management system (116). In some embodiments, the remote AV system (114) is involved in the installation of some or all of the vehicle's components, including an autonomous driving system, an autonomous vehicle computer, and software implemented by the autonomous vehicle computer. In some embodiments, the remote AV system (114) maintains (e.g., updates and / or replacements) such components and / or software throughout the vehicle's lifespan.

[0024] The fleet management system (116) includes at least one device configured to communicate with vehicles (102), a V2I device (110), a remote AV system (114), and / or a V2I infrastructure system (118). In one example, the fleet management system (116) includes a server, a group of servers, and / or other similar devices. In some embodiments, the fleet management system (116) is associated with a ridesharing company (e.g., an organization controlling the operation of a number of vehicles (e.g., vehicles including autonomous driving systems and / or vehicles not including autonomous driving systems)).

[0025] In some embodiments, the V2I system (118) includes at least one device configured to communicate with vehicles (102), a V2I device (110), a remote AV system (114), and / or a fleet management system (116) via a network (112). In some embodiments, the V2I system (118) is configured to communicate with the V2I device (110) via a connection different from the network (112). In some embodiments, the V2I system (118) includes a server, a group of servers, and / or other similar devices. In some embodiments, the V2I system (118) is associated with a municipal or private organization (e.g., a private organization maintaining the V2I device (110), etc.).

[0026] The number and arrangement of elements exemplified in FIG. 1 are provided as examples. There may be additional elements, fewer elements, different elements, and / or differently arranged elements than those exemplified in FIG. 1. Additionally or alternatively, at least one element of the environment (100) may perform one or more functions described as being performed by at least one different element of FIG. 1. Additionally or alternatively, at least one set of elements of the environment (100) may perform one or more functions described as being performed by at least one set of different elements of the environment (100).

[0027] Now, referring to FIG. 2, the vehicle (200) includes an autonomous driving system (202), a powertrain control system (204), a steering control system (206), and a brake system (208). In some embodiments, the vehicle (200) is identical or similar to the vehicle (102) (see FIG. 1). In some embodiments, the vehicle (102) has autonomous driving capabilities (e.g., fully autonomous vehicles (e.g., vehicles that do not rely on human intervention), highly autonomous vehicles (e.g., vehicles that do not rely on human intervention in certain situations), etc., without limitation, and implements at least one function, feature, device, etc. that enables the vehicle (200) to operate partially or fully without human intervention). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference may be made to SAE International's standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entirety of which is incorporated by reference. In some embodiments, the vehicle (200) is associated with an autonomous driving fleet manager and / or a ride-sharing company.

[0028] The autonomous driving system (202) includes a sensor suite comprising one or more devices such as cameras (202a), LiDAR sensors (202b), radar sensors (202c), and microphones (202d). In some embodiments, the autonomous driving system (202) may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), mileage measuring sensors that generate data associated with an indication of the distance traveled by the vehicle (200), etc.). In some embodiments, the autonomous driving system (202) uses one or more devices included in the autonomous driving system (202) to generate data associated with the environment (100) described herein. Data generated by one or more devices of the autonomous driving system (202) may be used by one or more systems described herein to observe the environment (e.g., environment (100)) where the vehicle (200) is located. In some embodiments, the autonomous driving system (202) includes a communication device (202e), an autonomous driving vehicle computer (202f), and a drive-by-wire (DBW) system (202h).

[0029] The cameras (202a) include at least one device configured to communicate with a communication device (202e), an autonomous vehicle computer (202f), and / or a safety controller (202g) via a bus (e.g., a bus identical or similar to the bus (302) of FIG. 3). The cameras (202a) include at least one camera (e.g., a digital camera using a light sensor such as a charge-coupled device (CCD), a thermal camera, an infrared (IR) camera, an event camera, etc.) for capturing images including physical objects (e.g., cars, buses, curbs, people, etc.). In some embodiments, the camera (202a) generates camera data as output. In some examples, the camera (202a) generates camera data including image data associated with the image. In this example, the image data may specify at least one parameter corresponding to the image (e.g., image characteristics such as exposure, brightness, etc., image timestamp, etc.). In such examples, the images may be in a format (e.g., RAW, JPEG, PNG, etc.). In some embodiments, the camera (202a) includes a plurality of independent cameras configured on the vehicle (e.g., located on the vehicle) to capture images for stereopsis (stereo vision). In some examples, the camera (202a) includes a plurality of cameras, and these plurality of cameras generate image data and transmit the image data to an autonomous vehicle computer (202f) and / or a fleet management system (e.g., a fleet management system identical or similar to the fleet management system (116) of FIG. 1). In such examples, the autonomous vehicle computer (202f) determines the depth to one or more objects within the field of view of at least two of the plurality of cameras based on image data from at least two cameras.In some embodiments, the cameras (202a) are configured to capture images of objects within a certain distance (e.g., up to 100 meters, up to 1 kilometer, etc.) from the cameras (202a). Accordingly, the cameras (202a) include features such as sensors and lenses optimized to recognize objects at one or more distances from the cameras (202a).

[0030] In one embodiment, the camera (202a) comprises at least one camera configured to capture one or more images associated with one or more traffic lights, street signs and / or other physical objects that provide visual navigation information. In some embodiments, the camera (202a) generates traffic light data associated with one or more images. In some examples, the camera (202a) generates TLD data associated with one or more images that include a format (e.g., RAW, JPEG, PNG, etc.). In some embodiments, the camera (202a) generating TLD data differs from other systems described herein that include cameras in that the camera (202a) may include one or more cameras having a wide field of view (e.g., a wide-angle lens, a fisheye lens, a lens having a field of view of approximately 120 degrees or more, etc.) to generate images of as many physical objects as possible.

[0031] The LiDAR (Laser Detection and Ranging) sensors (202b) include at least one device configured to communicate with a communication device (202e), an autonomous vehicle computer (202f), and / or a safety controller (202g) via a bus (e.g., a bus identical or similar to the bus (302) of FIG. 3). The LiDAR sensors (202b) include a system configured to transmit light from a light emitter (e.g., a laser transmitter). The light emitted by the LiDAR sensors (202b) includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensors (202b) encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensors (202b). In some embodiments, the light emitted by the LiDAR sensors (202b) does not penetrate the physical objects to which the light encounters. The LiDAR sensors (202b) also include at least one light detector that detects light emitted from a light emitter after it encounters a physical object. In some embodiments, at least one data processing system associated with the LiDAR sensors (202b) generates an image (e.g., a point cloud, a combined point cloud, etc.) representing objects included in the field of view of the LiDAR sensors (202b). In some examples, at least one data processing system associated with the LiDAR sensors (202b) generates an image representing the boundaries of physical objects, surfaces of physical objects (e.g., the topology of the surfaces), etc. In such examples, the image is used to determine the boundaries of physical objects within the field of view of the LiDAR sensors (202b).

[0032] The radar (Radar, Radio Detection and Ranging) sensors (202c) include at least one device configured to communicate with a communication device (202e), an autonomous vehicle computer (202f), and / or a safety controller (202g) via a bus (e.g., a bus identical or similar to the bus (302) of FIG. 3). The radar sensors (202c) include a system configured to transmit radio waves (in a pulsed or continuous manner). The radio waves transmitted by the radar sensors (202c) include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves transmitted by the radar sensors (202c) encounter a physical object and are reflected back to the radar sensors (202c). In some embodiments, the radio waves transmitted by the radar sensors (202c) are not reflected by some objects. In some embodiments, at least one data processing system associated with the radar sensors (202c) generates signals representing objects included in the field of view of the radar sensors (202c). For example, at least one data processing system associated with the radar sensors (202c) generates an image representing the boundaries of physical objects, surfaces of physical objects (e.g., topology of surfaces), etc. In some examples, the image is used to determine the boundaries of physical objects within the field of view of the radar sensors (202c).

[0033] The microphones (202d) include at least one device configured to communicate with a communication device (202e), an autonomous vehicle computer (202f), and / or a safety controller (202g) via a bus (e.g., a bus identical or similar to the bus (302) of FIG. 3). The microphones (202d) include one or more microphones (e.g., array microphones, external microphones, etc.) that capture audio signals and generate data associated with (e.g., indicating) the audio signals. In some examples, the microphones (202d) include transducer devices and / or similar devices. In some embodiments, one or more systems described herein may receive data generated by the microphones (202d) and determine the location (e.g., distance, etc.) of an object relative to the vehicle (200) based on audio signals associated with the data.

[0034] The communication device (202e) includes at least one device configured to communicate with cameras (202a), LiDAR sensors (202b), radar sensors (202c), microphones (202d), an autonomous vehicle computer (202f), a safety controller (202g), and / or a DBW system (202h). For example, the communication device (202e) may include a device identical or similar to the communication interface (314) of FIG. 3. In some embodiments, the communication device (202e) includes a vehicle-to-vehicle (V2V) communication device (e.g., a device that enables wireless communication of data between vehicles).

[0035] The autonomous vehicle computer (202f) includes at least one device configured to communicate with cameras (202a), LiDAR sensors (202b), radar sensors (202c), microphones (202d), a communication device (202e), a safety controller (202g), and / or a DBW system (202h). In some examples, the autonomous vehicle computer (202f) includes devices such as a client device, a mobile device (e.g., a cellular phone, a tablet, etc.), a server (e.g., a computing device including one or more central processing units, graphics processing units, etc.). In some embodiments, the autonomous vehicle computer (202f) is identical or similar to the autonomous vehicle computer (400) described herein. Additionally or alternatively, in some embodiments, the autonomous vehicle computer (202f) is configured to communicate with an autonomous vehicle system (e.g., an autonomous vehicle system identical or similar to the remote AV system (114) of FIG. 1), a fleet management system (e.g., a fleet management system identical or similar to the fleet management system (116) of FIG. 1), a V2I device (e.g., a V2I device identical or similar to the V2I device (110) of FIG. 1), and / or a V2I system (e.g., a V2I system identical or similar to the V2I system (118) of FIG. 1).

[0036] The safety controller (202g) includes at least one device configured to communicate with cameras (202a), LiDAR sensors (202b), radar sensors (202c), microphones (202d), a communication device (202e), an autonomous vehicle computer (202f), and / or a DBW system (202h). In some examples, the safety controller (202g) includes one or more controllers (electric controllers, electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle (200) (e.g., a powertrain control system (204), a steering control system (206), a brake system (208), etc.). In some embodiments, the safety controller (202g) is configured to generate control signals that take precedence over (e.g., ignore) the control signals generated and / or transmitted by the autonomous vehicle computer (202f).

[0037] The DBW system (202h) includes at least one device configured to communicate with a communication device (202e) and / or an autonomous vehicle computer (202f). In some examples, the DBW system (202h) includes one or more controllers (e.g., electric controllers, electromechanical controllers, etc.) configured to generate and / or transmit control signals for operating one or more devices of the vehicle (200) (e.g., powertrain control system (204), steering control system (206), brake system (208), etc.). Additionally or alternatively, one or more controllers of the DBW system (202h) are configured to generate and / or transmit control signals for operating at least one different device of the vehicle (200) (e.g., turn signals, headlights, door locks, window wipers, etc.).

[0038] The powertrain control system (204) includes at least one device configured to communicate with the DBW system (202h). In some examples, the powertrain control system (204) includes at least one controller, actuator, etc. In some embodiments, the powertrain control system (204) receives control signals from the DBW system (202h), and the powertrain control system (204) causes the vehicle (200) to start moving forward, stop moving forward, start moving backward, stop moving backward, accelerate in one direction, decelerate in one direction, perform a left turn, perform a right turn, etc. In one example, the powertrain control system (204) increases, keeps the energy (e.g., fuel, electricity, etc.) supplied to the vehicle's motor, or decreases it, thereby causing at least one wheel of the vehicle (200) to rotate or not rotate.

[0039] The steering control system (206) includes at least one device configured to rotate one or more wheels of the vehicle (200). In some examples, the steering control system (206) includes at least one controller, actuator, etc. In some embodiments, the steering control system (206) causes the two front wheels and / or two rear wheels of the vehicle (200) to rotate to the left or right so that the vehicle (200) turns to the left or right.

[0040] The brake system (208) includes at least one device configured to actuate one or more brakes to cause the vehicle (200) to reduce speed and / or remain at a stop. In some examples, the brake system (208) includes at least one controller and / or actuator configured to close one or more calipers associated with one or more wheels of the vehicle (200) on the corresponding rotor of the vehicle (200). Additionally or alternatively, in some examples, the brake system (208) includes an automatic emergency braking (AEB) system, a regenerative braking system, etc.

[0041] In some embodiments, the vehicle (200) includes at least one platform sensor (not explicitly exemplified) for measuring or inferring attributes of the state or condition of the vehicle (200). In some embodiments, the vehicle (200) includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), a wheel speed sensor, a wheel brake pressure sensor, a wheel torque sensor, an engine torque sensor, a steering angle sensor, etc.

[0042] Now, referring to FIG. 3, a schematic diagram of a device (300) is illustrated. As illustrated, the device (300) includes a processor (304), memory (306), a storage component (308), an input interface (310), an output interface (312), a communication interface (314), and a bus (302). In some embodiments, the device (300) corresponds to at least one device of the vehicles (102) (e.g., at least one device of the system of the vehicles (102)) and / or one or more devices of the network (112) (e.g., one or more devices of the system of the network (112)). In some embodiments, one or more devices of the vehicles (102) (e.g., one or more devices of the system of the vehicles (102)) and / or one or more devices of the network (112) (e.g., one or more devices of the system of the network (112)) include at least one device (300) and / or at least one component of the device (300). As illustrated in FIG. 3, the device (300) includes a bus (302), a processor (304), a memory (306), a storage component (308), an input interface (310), an output interface (312), and a communication interface (314).

[0043] The bus (302) includes a component that enables communication between components of the device (300). In some embodiments, the processor (304) is implemented in hardware, software, or a combination of hardware and software. In some examples, the processor (304) includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an acceleration processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component (e.g., a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.) that can be programmed to perform at least one function. The memory (306) includes random access memory (RAM), read-only memory (ROM), and / or other types of dynamic and / or static storage devices (e.g., flash memory, magnetic memory, optical memory, etc.) that store data and / or instructions to be used by the processor (304).

[0044] The storage component (308) stores data and / or software related to the operation and use of the device (300). In some examples, the storage component (308) includes a hard disk (e.g., magnetic disk, optical disk, magneto-optical disk, solid-state disk, etc.), a CD (compact disc), a DVD (digital versatile disc), a floppy disk, a cartridge, a magnetic tape, a CD-ROM, RAM, PROM, EPROM, FLASH-EPROM, NV-RAM, and / or other types of computer-readable media, together with a corresponding drive.

[0045] The input interface (310) includes a component that enables the device (300) to receive information, for example, through user input (e.g., a touchscreen display, keyboard, keypad, mouse, button, switch, microphone, camera, etc.). Additionally or alternatively, in some embodiments, the input interface (310) includes a sensor that detects information (e.g., a GPS (global positioning system) receiver, accelerometer, gyroscope, actuator, etc.). The output interface (312) includes a component that provides output information from the device (300) (e.g., a display, speaker, one or more light-emitting diodes (LEDs), etc.).

[0046] In some embodiments, the communication interface (314) includes a transceiver-like component (e.g., a transceiver, an individual receiver and a transmitter, etc.) that enables the device (300) to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some embodiments, the communication interface (314) enables the device (300) to receive information from another device and / or provide information to another device. In some embodiments, the communication interface (314) includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi® interface, a cellular network interface, etc.

[0047] In some embodiments, the device (300) performs one or more processes described herein. The device (300) performs these processes based on a processor (304) executing software instructions stored by a computer-readable medium, such as a memory (305) and / or a storage component (308). A computer-readable medium (e.g., a non-transient computer-readable medium) is defined herein as a non-transient memory device. A non-transient memory device includes a memory space located within a single physical storage device or a memory space distributed across multiple physical storage devices.

[0048] In some embodiments, software instructions are read into memory (306) and / or storage component (308) from another computer-readable medium or another device via a communication interface (314). When executed, the software instructions stored in memory (306) and / or storage component (308) cause the processor (304) to perform one or more processes described herein. Additionally or alternatively, hardwired circuits are used instead of or in combination with software instructions to perform one or more processes described herein. Accordingly, the embodiments described herein are not limited to any specific combination of hardware circuits and software unless otherwise explicitly stated.

[0049] The memory (306) and / or storage component (308) includes data storage or at least one data structure (e.g., a database, etc.). The device (300) may receive information from the data storage or at least one data structure within the memory (306) or storage component (308), store information therein, communicate information to it, or retrieve information stored therein. In some examples, the information includes network data, input data, output data, or any combination thereof.

[0050] In some embodiments, the device (300) is configured to execute software instructions stored in memory (306) and / or in the memory of another device (e.g., another device identical or similar to the device (300)). As used herein, the term “module” refers to at least one instruction stored in memory (306) and / or in the memory of another device that causes the device (300) (e.g., at least one component of the device (300)) to perform one or more processes described herein when executed by the processor (304) and / or by the processor of another device (e.g., another device identical or similar to the device (300)). In some embodiments, the module is implemented in software, firmware, hardware, etc.

[0051] The number and arrangement of components illustrated in FIG. 3 are provided as examples. In some embodiments, the device (300) may include additional components, fewer components, different components, or differently arranged components than illustrated in FIG. 3. Additionally or alternatively, a set of components of the device (300) (e.g., one or more components) may perform one or more functions described as being performed by other components or other sets of components of the device (300).

[0052] Now, referring to FIG. 4a, an exemplary block diagram of an autonomous vehicle computer (400) (sometimes referred to as the “AV stack”) is illustrated. As illustrated, the autonomous vehicle computer (400) includes a perception system (402) (sometimes referred to as the perception module), a planning system (404) (sometimes referred to as the planning module), a localization system (406) (sometimes referred to as the localization module), a control system (408) (sometimes referred to as the control module), and a database (410). In some embodiments, the perception system (402), the planning system (404), the localization system (406), the control system (408), and the database (410) are included in and / or implemented in the autonomous vehicle navigation system of the vehicle (e.g., the autonomous vehicle computer (202f) of the vehicle (200)). Additionally or alternatively, in some embodiments, the perception system (402), planning system (404), localization system (406), control system (408), and database (410) are included in one or more standalone systems (e.g., one or more systems identical or similar to the autonomous vehicle computer (400), etc.). In some embodiments, the perception system (402), planning system (404), localization system (406), control system (408), and database (410) are included in one or more standalone systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, some and / or all of the systems included in the autonomous vehicle computer (400) are implemented as software (e.g., software instructions stored in memory), computer hardware (e.g., a microprocessor, a microcontroller, an application-specific integrated circuit (ASIC), a Field Programmable Gate Array (FPGA), etc.), or a combination of computer software and computer hardware.It will also be understood that in some embodiments, the autonomous vehicle computer (400) is configured to communicate with a remote system (e.g., an autonomous vehicle system identical or similar to the remote AV system (114), a fleet management system identical or similar to the fleet management system (116), a V2I system identical or similar to the V2I system (118), etc.).

[0053] In some embodiments, the perception system (402) receives data associated with at least one physical object in the environment (e.g., data used by the perception system (402) to detect at least one physical object) and classifies at least one physical object. In some embodiments, the perception system (402) receives image data captured by at least one camera (e.g., cameras (202a)), and the image is associated with (e.g., represents) one or more physical objects within the field of view of at least one camera. In such examples, the perception system (402) classifies at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, pedestrians, etc.). In some embodiments, based on the classification of physical objects by the perception system (402), the perception system (402) transmits data associated with the classification of physical objects to the planning system (404).

[0054] In some embodiments, the planning system (404) receives data associated with a destination and generates data associated with at least one route (e.g., routes (106)) through which a vehicle (e.g., vehicles (102)) can travel toward the destination. In some embodiments, the planning system (404) receives data from the perception system (402) periodically or continuously (e.g., data associated with the classification of physical objects described above), and the planning system (404) updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system (402). In some embodiments, the planning system (404) receives data associated with the updated location of the vehicle (e.g., vehicles (102)) from the localization system (406), and the planning system (404) updates at least one trajectory or generates at least one different trajectory based on the data generated by the localization system (406).

[0055] In some embodiments, the localization system (406) receives data associated with (e.g., indicating therein) a location of a vehicle (e.g., vehicles (102)) in a region. In some embodiments, the localization system (406) receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensors (202b)). In certain embodiments, the localization system (406) receives data associated with at least one point cloud from a plurality of LiDAR sensors, and the localization system (406) generates a combined point cloud based on each of the point clouds. In these embodiments, the localization system (406) compares at least one point cloud or the combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the corresponding region stored in a database (410). Based on the localization system (406) comparing at least one point cloud or combined point cloud with a map, the localization system (406) then determines the location of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated prior to the operation of the vehicle. In some embodiments, the map includes, without limitation, a high-precision map of road geometric characteristics, a map describing road network connectivity characteristics, a map describing road physical characteristics (e.g., traffic speed, traffic volume, number of vehicle traffic lanes and cyclist traffic lanes, lane width, lane traffic direction, or lane marker type and location, or a combination thereof), and a map describing the spatial locations of road features, such as crosswalks, traffic signs, or various types of other driving signals. In some embodiments, the map is generated in real time based on data received by a perception system.

[0056] In another example, the localization system (406) receives GNSS (Global Navigation Satellite System) data generated by a GPS (global positioning system) receiver. In some examples, the localization system (406) receives GNSS data associated with the position of a vehicle within a corresponding area, and the localization system (406) determines the latitude and longitude of the vehicle within that area. In such examples, the localization system (406) determines the position of the vehicle within the corresponding area based on the latitude and longitude of the vehicle. In some embodiments, the localization system (406) generates data associated with the position of the vehicle. In some examples, the localization system (406) generates data associated with the position of the vehicle based on the determination of the position of the vehicle by the localization system (406). In such examples, the data associated with the position of the vehicle includes data associated with one or more semantic characteristics corresponding to the position of the vehicle.

[0057] In some embodiments, the control system (408) receives data associated with at least one trajectory from the planning system (404), and the control system (408) controls the operation of the vehicle. In some examples, the control system (408) receives data associated with at least one trajectory from the planning system (404), and the control system (408) controls the operation of the vehicle by generating and transmitting control signals that cause the powertrain control system (e.g., DBW system (202h), powertrain control system (204), etc.), the steering control system (e.g., steering control system (206)), and / or the brake system (e.g., brake system (208)) to operate. In an example where the trajectory includes a left turn, the control system (408) transmits a control signal that causes the vehicle (200) to turn left by causing the steering control system (206) to adjust the steering angle of the vehicle (200). Additionally or alternatively, the control system (408) generates and transmits control signals that cause other devices of the vehicle (200) (e.g., headlights, turn signals, door locks, window wipers, etc.) to change their states.

[0058] In some embodiments, the cognitive system (402), the planning system (404), the localization system (406), and / or the control system (408) implement at least one machine learning model (e.g., at least one multilayer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, at least one transformer, etc.). In some embodiments, the cognitive system (402), the planning system (404), the localization system (406), and / or the control system (408) implement at least one machine learning model alone or in combination with one or more of the systems mentioned above. In some embodiments, the cognitive system (402), the planning system (404), the localization system (406), and / or the control system (408) implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in an environment, etc.). Examples of machine learning model implementations are included below in relation to Figures 4b through 4d.

[0059] The database (410) stores data that is transmitted to and received from and / or updated by the perception system (402), planning system (404), localization system (406) and / or control system (408). In some examples, the database (410) includes a storage component (e.g., a storage component identical or similar to the storage component (308) of FIG. 3) that stores data and / or software related to operation and uses at least one system of the autonomous vehicle computer (400). In some embodiments, the database (410) stores data associated with a 2D and / or 3D map of at least one area. In some examples, the database (410) stores data associated with a 2D and / or 3D map of a part of a city, a part of a number of cities, a number of cities, a county, a state, a state (e.g., a country). In such an example, a vehicle (e.g., a vehicle identical or similar to vehicles (102) and / or a vehicle (200)) can be driven along one or more drivable areas (e.g., a single-lane road, a multi-lane road, a main road, a back road, an off-road trail, etc.) and at least one LiDAR sensor (e.g., a LiDAR sensor identical or similar to LiDAR sensors (202b)) can be made to generate data associated with an image representing objects included in the field of view of at least one LiDAR sensor.

[0060] In some embodiments, the database (410) is implemented across multiple devices. In some examples, the database (410) may be included in a vehicle (e.g., a vehicle identical or similar to vehicles (102) and / or vehicle (200)), an autonomous vehicle system (e.g., an autonomous vehicle system identical or similar to remote AV system (114)), a fleet management system (e.g., a fleet management system identical or similar to the fleet management system (116) of FIG. 1), a V2I system (e.g., a V2I system identical or similar to the V2I system (118) of FIG. 1), etc.

[0061] Now, referring to FIG. 4b, a diagram of the implementation of a machine learning model is illustrated. More specifically, a diagram of the implementation of a convolutional neural network (CNN) (420) is illustrated. For the sake of illustration, the following description of the CNN (420) will be made in relation to the implementation of the CNN (420) by the cognitive system (402). However, it will be understood that in some examples, the CNN (420) (e.g., one or more components of the CNN (420)) is implemented by systems different from or other than the cognitive system (402), such as the planning system (404), the localization system (406), and / or the control system (408). Although the CNN (420) includes certain features as described herein, these features are provided for illustrative purposes and are not intended to limit the present disclosure.

[0062] The CNN (420) includes a plurality of convolution layers, including a first convolution layer (422), a second convolution layer (424), and a convolution layer (426). In some embodiments, the CNN (420) includes a subsampling layer (428) (sometimes referred to as a pooling layer). In some embodiments, the subsampling layer (428) and / or other subsampling layers have a dimension (i.e., a quantity of nodes) smaller than the dimension of the upstream system. By having the subsampling layer (428) have a dimension smaller than the dimension of the upstream layer, the CNN (420) incorporates the amount of data associated with the initial input and / or the output of the upstream layer, thereby reducing the amount of computation required for the CNN (420) to perform downstream convolution operations. Additionally or alternatively, by having a subsampling layer (428) associated with at least one subsampling function (e.g., configured to do so) (as described below in relation to FIGS. 4c and 4d), the CNN (420) consolidates the amount of data associated with the initial input.

[0063] The cognitive system (402) performs convolution operations based on generating respective outputs by providing respective inputs and / or outputs associated with each of the first convolution layer (422), the second convolution layer (424), and the convolution layer (426). In some examples, the cognitive system (402) implements a CNN (420) based on providing data as inputs to the first convolution layer (422), the second convolution layer (424), and the convolution layer (426). In such an example, based on the perception system (402) receiving data from one or more different systems (e.g., one or more systems of a vehicle identical or similar to the vehicle (102), a remote AV system identical or similar to the remote AV system (114), a fleet management system identical or similar to the fleet management system (116), a V2I system identical or similar to the V2I system (118), etc., the perception system (402) provides the data as input to a first convolution layer (422), a second convolution layer (424), and a convolution layer (426). A detailed description of the convolution operations is included below in relation to FIG. 4c.

[0064] In some embodiments, the perception system (402) provides data associated with an input (referred to as the initial input) to the first convolution layer (422), and the perception system (402) uses the first convolution layer (422) to generate data associated with an output. In some embodiments, the perception system (402) provides the output generated by the convolution layer as an input to a different convolution layer. For example, the perception system (402) provides the output of the first convolution layer (422) as an input to a subsampling layer (428), a second convolution layer (424), and / or a convolution layer (426). In such an example, the first convolution layer (422) is referred to as the upstream layer, and the subsampling layer (428), the second convolution layer (424), and / or a convolution layer (426) are referred to as downstream layers. Similarly, in some embodiments, the perception system (402) provides the output of the subsampling layer (428) to the second convolution layer (424) and / or the convolution layer (426), and in this example, the subsampling layer (428) will be referred to as the upstream layer, and the second convolution layer (424) and / or the convolution layer (426) will be referred to as the downstream layers.

[0065] In some embodiments, before the cognitive system (402) provides input to the CNN (420), the cognitive system (402) processes data associated with the input provided to the CNN (420). For example, based on the cognitive system (402) normalizing sensor data (e.g., image data, LiDAR data, radar data, etc.), the cognitive system (402) processes data associated with the input provided to the CNN (420).

[0066] In some embodiments, the CNN (420) generates an output based on the cognitive system (402) performing convolution operations associated with each convolution layer. In some examples, the CNN (420) generates an output based on the cognitive system (402) performing convolution operations associated with each convolution layer and initial data. In some embodiments, the cognitive system (402) generates an output and provides the output as a fully connected layer (430). In some examples, the cognitive system (402) provides the output of the convolution layer (426) as a fully connected layer (430), wherein the fully connected layer (430) includes data associated with a plurality of feature values ​​referred to as F1, F2... FN. In this example, the output of the convolution layer (426) includes data associated with a plurality of output feature values ​​representing a prediction.

[0067] In some embodiments, the cognitive system (402) identifies a prediction among a plurality of predictions based on identifying a feature value associated with the prediction most likely to be accurate among a plurality of predictions. For example, if a fully connected layer (430) includes feature values ​​F1, F2, ... FN, and F1 is the largest feature value, the cognitive system (402) identifies the prediction associated with F1 as the accurate prediction among a plurality of predictions. In some embodiments, the cognitive system (402) trains the CNN (420) to generate predictions. In some examples, the cognitive system (402) trains the CNN (420) to generate predictions based on providing training data associated with predictions to the CNN (420).

[0068] Now, referring to FIGS. 4c and FIGS. 4d, a diagram illustrating the exemplary operation of a CNN (440) by a cognitive system (402) is illustrated. In some embodiments, the CNN (440) (e.g., one or more components of the CNN (440)) is identical or similar to the CNN (420) (e.g., one or more components of the CNN (420)) (see FIG. 4b).

[0069] In step (450), the perception system (402) provides data associated with an image as input to the CNN (440) (step (450)). For example, as illustrated, the perception system (402) provides data associated with an image to the CNN (440), where the image is a grayscale image represented as values ​​stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image, and the color image is represented as values ​​stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image, a radar image, etc.

[0070] In step (455), the CNN (440) performs a first convolution function. For example, the CNN (440) performs the first convolution function based on providing values ​​representing an image as input to one or more neurons (not explicitly exemplified) included in the first convolution layer (442). In this example, the values ​​representing the image may correspond to values ​​representing a region of the image (sometimes referred to as a receptive field). In some embodiments, each neuron is associated with a filter (not explicitly exemplified). The filter (sometimes referred to as a kernel) may be represented as an array of values ​​whose size corresponds to the values ​​provided as input to the neuron. In one example, the filter may be configured to identify edges (e.g., horizontal lines, vertical lines, straight lines, etc.). In successive convolutional layers, filters associated with neurons can be configured to identify successively more complex patterns (e.g., arcs, objects, etc.).

[0071] In some embodiments, the CNN (440) performs a first convolution function based on multiplying the values ​​provided as inputs for each of one or more neurons included in the first convolution layer (442) by the values ​​of the filter corresponding to each of one or more neurons. For example, the CNN (440) may generate a single value or an array of values ​​as output by multiplying the values ​​provided as inputs for each of one or more neurons included in the first convolution layer (442) by the values ​​of the filter corresponding to each of one or more neurons. In some embodiments, the collective output of the neurons of the first convolution layer (442) is referred to as the convolved output. In some embodiments, where each neuron has the same filter, the convolved output is referred to as the feature map.

[0072] In some embodiments, the CNN (440) provides the outputs of each neuron of the first convolution layer (442) to the neurons of the downstream layer. For clarity, the upstream layer may be a layer that transmits data to a different layer (referred to as the downstream layer). For example, the CNN (440) may provide the outputs of each neuron of the first convolution layer (442) to the corresponding neurons of the subsampling layer. In one example, the CNN (440) provides the outputs of each neuron of the first convolution layer (442) to the corresponding neurons of the first subsampling layer (444). In some embodiments, the CNN (440) adds a bias value to the aggregates of all values ​​provided to each neuron of the downstream layer. For example, the CNN (440) adds a bias value to the aggregates of all values ​​provided to each neuron of the first subsampling layer (444). In such an example, based on the aggregates of all values ​​provided to each neuron and the activation function associated with each neuron of the first subsampling layer (444), the CNN (440) determines the final value to be provided to each neuron of the first subsampling layer (444).

[0073] In step (460), the CNN (440) performs a first subsampling function. For example, the CNN (440) may perform the first subsampling function based on the CNN (440) providing the values ​​output by the first convolution layer (442) to the corresponding neurons of the first subsampling layer (444). In some embodiments, the CNN (440) performs the first subsampling function based on an aggregation function. In one example, the CNN (440) performs the first subsampling function based on the CNN (440) determining the maximum input among the values ​​provided to a given neuron (referred to as the max pooling function). In another example, the CNN (440) performs the first subsampling function based on the CNN (440) determining the average input among the values ​​provided to a given neuron (referred to as the average pooling function). In some embodiments, based on CNN (440) providing values ​​to each neuron of the first subsampling layer (444), CNN (440) generates an output, which is sometimes referred to as a subsampled convolved output.

[0074] In step (465), the CNN (440) performs a second convolution function. In some embodiments, the CNN (440) performs the second convolution function in a manner similar to how the CNN (440) performed the first convolution function as described above. In some embodiments, the CNN (440) performs the second convolution function based on providing the values ​​output by the first subsampling layer (444) as inputs to one or more neurons (not explicitly exemplified) included in the second convolution layer (446). In some embodiments, as described above, each neuron of the second convolution layer (446) is associated with a filter. As described above, the filter(s) associated with the second convolution layer (446) may be configured to identify more complex patterns than the filter associated with the first convolution layer (442).

[0075] In some embodiments, the CNN (440) performs a second convolution function based on multiplying the values ​​provided as inputs for each of one or more neurons included in the second convolution layer (446) by the values ​​of the filter corresponding to each of one or more neurons. For example, the CNN (440) may generate a single value or an array of values ​​as output by multiplying the values ​​provided as inputs for each of one or more neurons included in the second convolution layer (446) by the values ​​of the filter corresponding to each of one or more neurons.

[0076] In some embodiments, the CNN (440) provides the outputs of each neuron of the second convolution layer (446) to the neurons of the downstream layer. For example, the CNN (440) may provide the outputs of each neuron of the first convolution layer (442) to the corresponding neurons of the subsampling layer. In one example, the CNN (440) provides the outputs of each neuron of the first convolution layer (442) to the corresponding neurons of the second subsampling layer (448). In some embodiments, the CNN (440) adds a bias value to the aggregates of all values ​​provided to each neuron of the downstream layer. For example, the CNN (440) adds a bias value to the aggregates of all values ​​provided to each neuron of the second subsampling layer (448). In such an example, based on the aggregates of all values ​​provided to each neuron and the activation function associated with each neuron of the second subsampling layer (448), the CNN (440) determines the final value to be provided to each neuron of the second subsampling layer (448).

[0077] In step (470), the CNN (440) performs a second subsampling function. For example, the CNN (440) may perform the second subsampling function based on the CNN (440) providing values ​​output by the second convolution layer (446) to the corresponding neurons of the second subsampling layer (448). In some embodiments, the CNN (440) performs the second subsampling function based on the CNN (440) using an aggregation function. In one example, as described above, the CNN (440) performs the first subsampling function based on the CNN (440) determining the maximum input or average input among the values ​​provided to a given neuron. In some embodiments, the CNN (440) generates an output based on the CNN (440) providing values ​​to each neuron of the second subsampling layer (448).

[0078] In step (475), the CNN (440) provides the output of each neuron of the second subsampling layer (448) to the fully connected layers (449). For example, the CNN (440) provides the output of each neuron of the second subsampling layer (448) to the fully connected layers (449) so that the fully connected layers (449) generate an output. In some embodiments, the fully connected layers (449) are configured to generate an output associated with a prediction (sometimes referred to as a classification). The prediction may include an indication that an object included in an image provided as input to the CNN (440) includes an object, a set of objects, etc. In some embodiments, the cognitive system (402) performs one or more operations as described herein and / or provides data associated with the prediction to a different system.

[0079] FIG. 5 is a block diagram illustrating an exemplary process (500) of agent trajectory prediction. In some embodiments, the process (500) is implemented using an autonomous driving system identical or similar to the autonomous driving system (202) described with reference to FIG. 2 (e.g., entirely, partially, etc.). In some embodiments, one or more steps of the process (500) are performed by another device or system, or by another group of devices and / or systems that are separate from or include the autonomous driving system (e.g., entirely, partially, etc.). For example, one or more steps of the process (500) may be performed by a remote AV system (114), a vehicle (102) (e.g., the autonomous driving system (202) of the vehicle (102 or 200)), the device (300) of FIG. 3 and / or an AV computer (400) (e.g., one or more systems of the AV computer (400) of FIG. 4a) (e.g., entirely, partially, etc.). In some embodiments, the steps of the process (500) may be performed in cooperation between any of the systems mentioned above.

[0080] In some embodiments, high definition (HD) maps and agent tracks (e.g., previous agent trajectories) are encoded using a graph representation of the scene. The graph structure of the scene is utilized to explicitly model the lateral or root variability of the trajectories. However, instead of aggregating the entire scene context into a single vector and learning a one-to-many mapping for multiple trajectories, the prediction of the trajectories is conditioned on the context optionally aggregated based on paths traversed in the directed graph by individual policies. For example, the context along a sampled path (502) is aggregated, and the trajectory (504) is predicted based on the context of the sampled path (502). In some embodiments, the policy is a state-action mapping that mimics the expert's action (ground truth) in the optionally aggregated context.

[0081] In some embodiments, a portion of the scene context (e.g., context following the sampled path (502)) is optionally aggregated for each prediction by sampling the path traversal from a learned behavior cloning policy (the policy is learned by the behavior cloning) (e.g., the sampled path (502)).

[0082] By directly selecting a subset of graphs (sampled paths) used for each prediction, the representation requirements for the trajectory decoder can be reduced. Additionally, the stochastic policy generates various sets of sampled paths and captures the lateral variability of the multimodal trajectory distribution.

[0083] In some embodiments, the prediction is additionally conditioned on a sampled latent variable to account for longitudinal variability of the trajectory. This enables the prediction of individual trajectories even for the same path trajectory.

[0084] FIG. 6 is a block diagram illustrating an exemplary agent trajectory prediction system (600). In some embodiments, the agent trajectory prediction system (600) is included in the autonomous driving system (202) of FIG. 2, the device (300) of FIG. 3, or the autonomous driving vehicle computer (400) of FIG. 4a. An exemplary agent trajectory prediction system (600) includes a graph encoder (602), a policy header (604), and a trajectory decoder (606). The graph encoder (602) encodes the agent context and map context (agent features and map features) as node encodings of a directed graph. In an example, the agent context characterizes the respective attributes associated with each agent, such as agent behavior and physical features. In some embodiments, the agent context also characterizes the relationship of each agent with other agents in a scene (e.g., social context). In the example, the map context characterizes static attributes of the environment, such as drivable areas, non-drivable areas, and traffic control elements (e.g., traffic lights, stop signs).

[0085] The policy header (604) learns a policy for graph traversal. The policy maps the states in the directed graph to a probability distribution for actions and then generates a distribution for graph traversal. The policy is trained based on how an expert (real-world data) traversed the graph. The trajectory decoder (606) predicts a set of trajectories for an agent (e.g., a target vehicle in the scene) based on node encoding and sampled latent variables according to the policy learned by the policy header (604). In some examples, the target vehicle is the agent for which the predicted set of trajectories is generated. The graph encoder (602) further includes a lane graph generator (608), a gated recurrent unit (GRU) encoder (610), an agent node attention layer (612), and a graph neural network (GNN) layer (614). As illustrated in FIG. 6, a graph having a final node encoding is generated by a graph encoder (602) and input to a policy header (604) and a trajectory decoder (606). The graph encoder (602) is described in detail in FIG. 7, which further illustrates an exemplary graph encoder (602). The policy header (604) is described in detail in FIG. 9, which further illustrates an exemplary policy header (604). The trajectory decoder (606) is described in detail in FIG. 10, which further illustrates an exemplary trajectory decoder (606).

[0086] The future trajectory of an agent in a scene is predicted based on the agent's respective past trajectory and the scene's HD map. In some embodiments, the future trajectory of a vehicle of interest (e.g., a subset of agents) is predicted based on the vehicle of interest's past trajectory, the past trajectories of nearby vehicles and pedestrians, and the scene's HD map. The scene is represented in a bird's-eye view, and predictions are made using the bird's-eye view representation of the scene. Agent-centered reference frames aligned along the agent's instantaneous motion direction are used for predictions associated with each agent. For example, referring to FIG. 5, the trajectory (504) is illustrated in an agent-centered reference frame aligned along the agent's instantaneous motion direction.

[0087] FIG. 7 illustrates an exemplary graph encoder (602). The graph encoder (602) of FIG. 7 corresponds to the graph encoder (602) of FIG. 6. In the graph encoder (602), the lane graph generator (608) generates a directed graph including nodes and edges based on a high-definition (HD) map of a driving scene. It generates. In some examples, one HD map corresponds to one directed graph. The graph structure generated by the graph encoder (602) is further described in relation to FIGS. 11a through 11c.

[0088] The HD map of the driving scene provides a representation of road terrain and traffic rules. FIG. 8 illustrates an exemplary HD map (802) and an exemplary direction graph (804) corresponding to the exemplary HD map (802). As illustrated in FIG. 8, the exemplary HD map (802) includes lanes (806) (represented as straight lines or polylines) and crosswalks and stop lines (808) (represented as polygons). Multiple agents or road actors (810) are driving in the lanes (806). A network of lane centerlines captures the direction of traffic flow and the legal paths or routes that each driver can follow. Lane centerlines are nodes ( It is represented as ). To ensure that each node represents a lane segment of similar length, longer lane centerlines are divided or subdivided into smaller segments of fixed length (e.g., the maximum length of each segment is 20 meters), and each segment is discretized into a set of N poses (e.g., each polyline is discretized at a resolution of 1 meter). In some embodiments, each segment corresponds to a node in a directed graph (804). Each node is a sequence of feature vectors It is expressed as, and each , n∈[1,N] and, here , and Is The position and yaw of the nth pose, and is a 2-D binary vector indicating whether the pose is placed on a stop line or a crosswalk. Therefore, the node feature captures not only geometric structures (e.g., lane outlines) along the lane centerline but also traffic control elements (e.g., crosswalks, stop lines).

[0089] In some embodiments, edges in the directed graph (804) are arranged so that any traversed path through the graph corresponds to a legal route that a vehicle can take in the scene. There are two types of edges: successor edges and proximal edges. Successor edges (E suc ) connects a node to the next node along a lane. In some examples, a follow-up edge connects a node to the next node along the same lane in the direction of travel. A given node can have multiple follow-up edges. For example, when a lane branches into multiple lanes at an intersection, a node has multiple follow-up edges. Similarly, multiple nodes can have the same follow-up edge. For example, when two or more lanes merge, multiple nodes have the same follow-up edge. Proximal edge (E prox ) connects a node in the first lane to the next node in the second lane. In this way, proximate edges are used to describe lane changes. In the example, proximate edges connect neighboring lane nodes when the nodes are within a distance threshold of each other and the motion direction associated with each node is within a yaw threshold. The yaw threshold ensures that proximate edges are not misassigned at intersections where multiple lanes intersect.

[0090] In some embodiments, a directional graph (804) is generated for all lane centerlines within a fixed area of ​​an HD map around a target vehicle (e.g., a vehicle of interest, such as one of a plurality of agents or road actors (810)). For example, map elements within an area of ​​[-50, 50] meters laterally and [-20, 80] meters longitudinally around the target vehicle may be used to generate the directional graph (804).

[0091] In some embodiments, the past trajectory of the agent (810) in the scene is obtained from an onboard detector and a multi-object tracker. In the example, the past trajectory of agent i is past t h A sequence of motion state vectors for a time step It is expressed as. Each and, here is the BEV (bird's-eye view) position coordinate, and , and is the agent's velocity, acceleration, and yaw rate at time t, and is an indicator (a value of 1 indicates a pedestrian, while a value of 0 indicates a vehicle).

[0092] Referring again to FIGS. 6 and 7, the GRU encoder (610) encodes agent features, node features, and motion features to generate node encoding (702), agent encoding (704), and motion encoding (706). Both the agent trajectory and the lane polyline form sequences of features in a defined order and are encoded independently using the GRU encoder (610). In some embodiments, the target vehicle trajectory , surrounding vehicle trajectories (Vehicles around the target vehicle) and initial node features Three GRU encoders (610) are provided to encode. Each of the three GRU encoders (610) is motion encoding (702), agent encoding (704) and initial node encoding Prints (706).

[0093] In some embodiments, possible node encodings include position in the agent-centered frame (e.g., the average (x,y) position of all points within a lane segment), motion direction (e.g., the (x,y) difference between the last and first points within a lane segment; curvature; centerline encoding using a 1D CNN; RNN, lane width / lateral distance to the road boundary, and other scene elements (e.g., crosswalk, stop line, road boundary). In some embodiments, a binary flag may be used to indicate whether a node is placed on a crosswalk. The binary flag may be used to indicate the adjacent sidewalk to the left and separately to the right. In some embodiments, if an off-road trajectory is considered, there may be additional absorption nodes at the road boundary where the policy ends.

[0094] In some embodiments, the encoder (610) is a Recurrent Neural Network (RNN) encoder, such as a Long Short-Term Memory Networks (LSTM) encoder.

[0095] The agent node attention layer (612) updates the initial node encoding. A driver (e.g., an agent) navigates through a traffic scene in cooperation with other drivers (e.g., agents) and pedestrians. Therefore, since surrounding agents can interact with the target vehicle's route or path, surrounding agents serve as useful clues for trajectory prediction. In some embodiments, the node encoding is updated to the agent encoding using scaled dot product attention. In the example, attention weights are calculated by normalizing the output score of a feed-forward neural network described by a function that captures the alignment between the input at j and the output at i. The agent node attention layer generates node embeddings that take into account the context of the surrounding agents. In the example, the agent node attention layer (612) calculates three vectors (query, key, and value) for each agent. To update the node encoding to the agent encoding, the dot product of the agent's query vector and the key vectors of all other agents is determined.

[0096] In some embodiments, agent encodings corresponding to nearby agents within the distance threshold of each node are considered in the update to the initial node encoding. This allows the trajectory decoder (606) to selectively focus on agents capable of interacting with the predicted trajectory of the target vehicle. Keys and values ​​are the encodings of nearby agents. It is obtained by linearly projecting it, and the query is It is obtained by linearly projecting it. In some embodiments, the updated node encoding is obtained by concatenating the output of the agent node attention layer (612) with the initial node encoding.

[0097] The GNN layer (614) aggregates local context from neighboring nodes and outputs a final node encoding. In some examples, the GNN layer (614) is a neural model (e.g., GNN) that captures the dependencies of the graph through messages passed between the nodes of the graph. In some examples, the GNN is an optimizable transformation for all properties of the graph (nodes, edges, and global context) that maintains graph symmetry (permutation invariant). The GNN takes as input a graph having an initial node encoding associated with node, edge, and context information. The encoding is transformed incrementally without changing the connectivity of the input graph. In this way, a directed graph having node encodings for all lane centerlines within a fixed area of ​​the HD map around the target vehicle is generated. In some embodiments, the GNN layer (614) is a graph convolution network (GCN) layer or a graph attention network (GAT) layer that aggregates local context from neighboring nodes. In some embodiments, agent context and map context are aggregated using a GCN network layer as described in relation to FIG. 12. Additionally, in some embodiments, subsequent edges and proximate edges are treated as equivalent and bidirectional, which enables the aggregation of context from all directions around each node. The output of the GNN layer (614) serves as the final node encoding learned by the graph encoder (602).

[0098] FIG. 9 illustrates an exemplary policy header (604). The exemplary policy header (604) is used to output a discrete probability distribution (908) for outgoing edges (e.g., subsequent edges, proximate edges) from each node, and enables sampling of paths in a directed graph (804) (Fig. 8), which are the most likely routes that the target vehicle will take in the future. All paths in the directed graph (804) (Fig. 8) correspond to plausible routes for the target vehicle. However, not all routes are equally likely to be the routes the target vehicle will take. For example, past motion of the target vehicle approaching an intersection may indicate that the target vehicle is preparing to turn rather than proceed straight. In the example, a slow-moving lane makes it more likely that the target vehicle will change lanes rather than remain in the same lane. For convenience of explanation, the policy is described as a discrete probability distribution. However, in some embodiments, the policy header (604) learns a policy for graph traversal according to a reward model (e.g., the reward model (1300) of FIG. 13).

[0099] The policy's sampled rollout corresponds to the routes or paths the target vehicle is likely to take in the future. This graph traversal is learned. In some embodiments, the rollout is the simulated result of a policy (e.g., state-action mapping) that generates trajectory samples. The policy is represented as a discrete probability distribution (908) for edges leaving each node. Policy Edges from all nodes to the final state are included so that termination can occur at this target location. Edge probabilities are output by the policy header (604). In the policy header (604), the final node encoding (902) and motion encoding (904) for the target vehicle are obtained from the directed graph (e.g., the directed graph (804) of FIG. 8). The policy header (604) includes a scoring layer (906). The scoring layer (906) for each edge It includes a multilayer perceptron (MLP) network with shared weights that outputs a scalar score for . Equation 1 for calculating the scalar score is provided below:

[0100] (1)

[0101] Therefore, the scoring function considers not only the motion of the target vehicle but also the agent context and local scene at specific edges. The score is normalized using a softmax layer for all outgoing edges at each node to output a policy for graph traversal. Equation 2 for outputting the policy distribution (908) is provided below:

[0102] (2)

[0103] In some embodiments, the policy header (604) is trained using behavior replication. For each prediction instance, the actual trajectory is used to determine which node the target vehicle has visited. In some embodiments, nodes whose motion direction is within the yaw threshold of the target vehicle's pose are considered in the current prediction. Edge It is treated as visited if both nodes u and v are visited. Measured trajectory( The negative log-likelihood of the edge probabilities for all edges visited by ) is used as the loss function to train the graph traversal policy. Equation 3 for outputting the loss function is provided below:

[0104] (3)

[0105] Referring again to FIG. 6, the trajectory decoder (606) is used to output a predicted trajectory conditioned on a path traversed in a directed graph output by a graph encoder according to a policy output by a policy header (604) and a sampled latent variable.

[0106] FIG. 10 illustrates an exemplary orbit decoder (606). Policy The sampling rollout yields a plausible future route or path for the target vehicle based on the final node encoding (1006). In some embodiments, the most relevant context for predicting the future trajectory follows these routes. An exemplary trajectory decoder (606) uses a multihead attention layer (1002) to optionally aggregate the context along the sampled route or path.

[0107] In some embodiments, a sequence of nodes corresponding to the sampled policy rollout (1008) Given, an exemplary trajectory decoder (606) uses multi-head scaled inner product attention to aggregate the map and agent context for the node sequence as illustrated in FIG. 10. The motion encoding of the target vehicle is projected linearly to obtain the query, and the node features It is projected linearly to obtain key and value for calculating attention. The multihead attention layer (1002) encodes the context along the sampled path traversed according to the policy using a context vector. Prints (1010). Each individual policy rollout is an individual context vector It calculates and enables the prediction of a trajectory along various sets of routes. The variety of routes alone cannot explain the multimodality of the future trajectory. The driver (e.g., agent) may brake, accelerate, and / or follow different motion profiles (e.g., position, velocity, and / or acceleration) along the planned route. In some embodiments, to enable the trajectory decoder (606) to output individual motion profiles, the trajectory prediction (1016) also includes sampled latent variables Conditioned on (1014). Unlike the root, vehicle speed and acceleration depend on the continuum. In some embodiments, latent variables (1014) is sampled from a continuous distribution. In some embodiments, latent variables (1014) is modeled as a multivariate standard normal distribution.

[0108] Trajectory In order to predict, The rollout of is sampled is obtained, and from the latent distribution This is sampled. and Is Connected to, , , is passed through MLP(1004) It prints, and this It is the future position (e.g., x and y coordinates) for a time step (e.g., predicted horizon). Equation 4 for outputting is provided below:

[0109] (4)

[0110] The sampling process can often be redundant and yields similar or repeating trajectories. However, lightweight encoder and decoder heads enable the sampling of a large number of trajectories in parallel.

[0111] of the trajectory distribution To obtain the final set of modes, K-means clustering is used to determine the cluster centers of the final set of predictions. Output as.

[0112] The trajectory decoder (606) is used to avoid penalizing various valid trajectories output by the trajectory decoder (606), and the actual trajectory ( It is trained using the winner-take-all mean displacement error for ). Equation 5 for outputting the loss function (e.g., regression loss) is provided below:

[0113] (5)

[0114] The trajectory decoder (606) is end-to-end trained using a multi-task loss that combines the losses from Equations 3 and 5. Equation 6 for combining the loss functions is provided below.

[0115] (6)

[0116] FIGS. 11a through 11c illustrate three exemplary directed graphs having graph structures generated by the lane graph generator (608) of FIGS. 6 and 7. In the example of FIG. 11a, the directed graph (1102) includes successor edges (1108), each successor edge connecting a node along a lane to the next node. Thus, the directed graph (1102) of FIG. 11a is structured with predecessor-successor connections between nodes. The directed graph structure (1102) of FIG. 11a does not include edges indicating a lane change or an illegal maneuver. In the example of FIG. 11b, the directed graph (1104) includes successor edges (1108B) and proximate edges (1110B) connecting adjacent nodes for a lane change. In the example of FIG. 11c, the directed graph (1106) includes a subsequent edge (1108C), a proximate edge (1110C), and an edge representing an illegal maneuver (1112C) (e.g., a route that does not comply with traffic laws). The directed graph (1106) connects all nearby nodes.

[0117] FIG. 12 illustrates the aggregation of agent contexts and map contexts using graph convolution. As illustrated in FIG. 12, an "Actor-to-Lane" context aggregation (1202) propagates real-time traffic information from agents to lane features. For example, real-time traffic information includes lanes occupied by agents. A "Lane-to-Lane" context aggregation (1204) propagates traffic information along valid paths in the lane graph. A "Lane-to-Actor" context aggregation (1206) fuses the latest lane information from the scene into the agents. An "Actor-to-Actor" context aggregation (1208) propagates real-time information between agents. In some embodiments, contexts are aggregated using a graph convolutional network (GCN) layer with extended convolution. Different edge types may encode directions with different learnable weights. In this way, the agent context is encoded into the node features. Extended graph convolution propagates features over longer distances along the lanes.

[0118] FIG. 13 shows the policy of the policy header (e.g., policy header (604)). An exemplary reward model (1300) for determining is illustrated. The reward model (1300) includes a state space and an action space (1303) (edges) containing nodes S (1302) and S' (1304). When an edge is selected, the next node is deterministic (e.g., deterministic dynamics (T(s, A)). The reward model (1300) is used to inhibit or encourage specific state transitions rather than specific states. The reward model can be expressed by Equation 7 below:

[0119] (7)

[0120] Here is a node encoding from GCN (1310). The next node S' (1302) is determined by applying an edge to the current node S (1302), or am. is a one-hot encoding of the edge type, and e() (1306) is a fully connected (FC) embedding layer used to obtain the edge type encoding (1312). The node encoding (1310), edge type encoding (1312), and destination node encoding (1314) of the source node are connected (1316). A network (1318) is applied to the connected node encoding (1310, 1314) and edge type encoding (1312). In some embodiments, the network (1318) includes a 1x1 convolution layer having weights shared for all nodes. The output is a reward value (1320) associated with the edge (1303).

[0121] During the training of the reward model (1300), the current reward is used with soft value iterations. V(s) and Q(s,a) for (1320) are estimated. In the example, the maximum entropy (MaxEnt) policy is used for training, and Equation 8 for the MaxEnt policy is provided below:

[0122] (8)

[0123] State-Action D W (s, a) The visiting frequency is p W It is calculated for (). Backpropagation is D GT (s, a) - D W It is calculated as (s, a), where D GT (s, a) is the measured state-action visit frequency. r W When (s, a) is obtained, during inference, the policy is determined, the rollout of the policy is sampled to traverse the graph, and the estimation of V(s) and Q(s,a) and the MaxEnt policy are repeated to assign probabilities.

[0124] As described above, when an edge is selected, the next node is deterministic (e.g., deterministic dynamic T(s, a)). In some embodiments, the dynamics are integrated into the reward model. For example, a second encoder-decoder model is used to generate a trajectory. In this example, for each node along the sampled path in the graph, the node features are encoded using the second encoder. The second decoder obtains the node encoding from the second encoder and outputs the velocity along and across the centerline over time. In the example, the velocity is output in Frenet frames for the centerline.

[0125] In some embodiments, a fixed motion profile is used to determine nodes and edges for generating a trajectory when determining the next node using a reward model. Using a fixed motion profile, lane centerlines are connected for the sampled path in the graph. A smoothing spline is used for adjacent or illegitimate edges to determine the centerline. A constant velocity or constant acceleration motion profile is used to obtain the trajectory.

[0126] In some embodiments, in the reward model, individual acceleration values ​​are added to each action (e.g., edge) and individual velocities are added to each state (e.g., node). Instead of a fixed motion profile, K individual acceleration values ​​are included in the action. In an embodiment, the reward model (1300) outputs K values ​​for each node feature / acceleration value pair. The acceleration values ​​are sampled with the next node according to a policy, and the acceleration values ​​are used to predict the trajectory along the centerline.

[0127] FIG. 14 illustrates an exemplary process (1400) for predicting a trajectory for an agent. In some embodiments, the process (1400) is implemented using an autonomous driving system identical or similar to the autonomous driving system (202) described with reference to FIG. 2 (e.g., entirely, partially, etc.). In some embodiments, one or more steps of the process (1400) are performed by another device or system, or by another group of devices and / or systems that are separate from or include the autonomous driving system (e.g., entirely, partially, etc.). For example, one or more steps of the process (1400) may be performed by a remote AV system (114), a vehicle (102) (e.g., the autonomous driving system (202) of the vehicle (102 or 200)), the device (300) of FIG. 3 and / or an AV computer (400) (e.g., one or more systems of the AV computer (400) of FIG. 4a) (e.g., entirely, partially, etc.). In some embodiments, the steps of the process (1400) may be performed in cooperation between any of the systems mentioned above.

[0128] In block 1402, a directed graph is generated by encoding map features and agent features as node encodings for the directed graph. In some embodiments, the graph includes learned representations for nodes in the graph that include map features and agent features. Generating the graph involves encoding map features and agent features as node encodings for each node of the graph. The directed graph includes learned representations for nodes in the graph. The learned representations include map features and agent features.

[0129] In block 1404, a policy is determined. A policy is determined to be applied to the outgoing edges of the nodes in the graph. In some embodiments, the policy is a discrete probability distribution, which is a probability distribution for the outgoing edges from each node for the target vehicle. In some embodiments, the policy is based on a reward model. In block 1406, one or more paths are sampled for the target vehicle in the scene according to the policy. In block 1408, a set of trajectories is predicted based on the sampled paths traversed by the policy and the sampled latent variables. In block 1410, the AV is operated based on the predicted set of trajectories for the target vehicle.

[0130] According to some non-limiting embodiments or examples, a method is provided comprising: generating a graph corresponding to a map of a scene by encoding map features and agent features as node encodings of the graph using at least one processor; determining a policy to apply to edges outgoing from nodes of the graph using at least one processor; sampling a path for a target vehicle in the scene according to the policy using at least one processor using at least one processor; predicting a set of trajectories based on the sampled path and sampled latent variables traversed by the policy using at least one processor using at least one processor; and operating a vehicle based on the set of trajectories of the target vehicle using at least one processor.

[0131] According to some non-limiting embodiments or examples, a system is provided comprising: a graph encoder for encoding high-resolution maps and agent features into a graph for generating a final node encoding, wherein the graph comprises nodes and edges, wherein the nodes represent segments of lane centerlines and the edges represent transitions between nodes, and the graph is used to generate the final node encoding; a policy header for learning a policy for traversing a sampled graph based on the motion of a target vehicle as well as agent context and local scenes at neighboring nodes; and a trajectory decoder for predicting a trajectory based on the node encoding and sampled latent variables along the path traversed by the policy.

[0132] According to some non-limiting embodiments or examples, at least one non-transient storage medium storing instructions is provided, wherein when the instructions are executed by at least one processor, the at least one processor generates a graph corresponding to a map of a scene by encoding map features and agent features as node encodings of the graph; determines a policy to apply to edges outgoing from nodes of the graph; samples a path for a target vehicle in the scene according to the policy; predicts a set of trajectories based on the sampled path and sampled latent variables traversed by the policy; and operates the vehicle based on the set of trajectories of the target vehicle.

[0133] Clause 1: A method comprising: generating a graph corresponding to a map of a scene by encoding map features and agent features as node encodings of the graph using at least one processor; determining a policy to apply to edges outgoing from nodes of the graph using at least one processor; sampling a path for a target vehicle in the scene according to the policy using at least one processor using at least one processor; predicting a set of trajectories based on the sampled path and sampled latent variables traversed by the policy using at least one processor using at least one processor; and operating a vehicle based on the set of trajectories of the target vehicle using at least one processor.

[0134] Clause 2: A method according to Clause 1, wherein each node corresponds to a segment of the lane centerline of the map.

[0135] Clause 3: A method comprising, in Clause 1 or 2, further a step of updating the node encoding to a surrounding agent encoding by calculating a scaled inner attention weight.

[0136] Clause 4: A method comprising the step of using a graph neural network to aggregate local context from neighboring nodes into the node encoding of the graph in any one of Clauses 1 to 3.

[0137] Clause 5: A method in which, in any one of Clauses 1 to 4, the policy for applying to the outgoing edge is a discrete probability distribution for the outgoing edge at a node of the graph.

[0138] Clause 6: A method in which, in any one of Clauses 1 through 5, the policy is predicted by training a multilayer perceptron (MLP) using behavior replication.

[0139] Clause 7: A method comprising, in any one of Clauses 1 to 6, optionally aggregating context along the sampled path and predicting the set of trajectories based on the sampled path traversed by the policy, the aggregated context, and the sampled latent variables.

[0140] Clause 8: The method of Clause 7, wherein the step of predicting the trajectory set comprises: connecting the aggregated context and the sampled latent variable to the motion encoding; and inputting the connected aggregated context and the sampled latent variable to a multilayer perceptron, wherein the trajectory set indicates a predicted position at a future time step.

[0141] Clause 9: A system comprising: a graph encoder for encoding high-resolution maps and agent features into a graph for generating a final node encoding, wherein the graph comprises nodes and edges, wherein the nodes represent segments of lane centerlines and the edges represent transitions between nodes, and the graph is used to generate the final node encoding; a policy header for learning a policy for traversing a sampled graph based on the motion of a target vehicle as well as agent context and local scenes at neighboring nodes; and a trajectory decoder for predicting a trajectory based on the node encoding and sampled latent variables along a path traversed by the policy.

[0142] Clause 10: A system in which, in Clause 9, the policy is a discrete probability distribution of transitions associated with each edge at each node.

[0143] Clause 11: A system according to Clause 9 or 10, wherein the graph encoder comprises one or more gate circulation units for encoding a target vehicle trajectory, surrounding vehicle trajectories, and node features.

[0144] Clause 12: A system in which, in any one of Clauses 9 to 11, the trajectory decoder comprises a multihead attention layer that outputs a context vector for each policy, and the context vector is combined with motion encoding and the sampled latent variable to predict the trajectory.

[0145] Clause 13: A system in which, in any one of Clauses 9 through 12, the initial node encoding is updated to a surrounding agent encoding by calculating scaled inner product attention weights to generate the final node encoding.

[0146] Clause 14: A system in which, in any one of Clauses 9 through 13, the graph encoder is configured to use a graph neural network to aggregate local context from neighboring nodes into the final node encoding of the graph.

[0147] Clause 15: At least one non-transient storage medium storing instructions, wherein when the instructions are executed by at least one processor, the at least one processor: generates a graph corresponding to a map of a scene by encoding map features and agent features as node encodings of the graph; determines a policy to apply to edges outgoing from nodes of the graph; samples a path for a target vehicle in the scene according to the policy; predicts a set of trajectories based on the sampled path and sampled latent variables traversed by the policy; and operates the vehicle based on the set of trajectories of the target vehicle.

[0148] Clause 16: At least one non-transient storage medium in which, in Clause 15, each node corresponds to a segment of the lane centerline of the map.

[0149] Clause 17: At least one non-transient storage medium comprising updating the node encoding to a surrounding agent encoding by calculating a scaled inner attention weight in Clause 15 or 16.

[0150] Clause 18: At least one non-transient storage medium comprising, in any one of Clauses 15 to 17, using a graph neural network to aggregate local context from neighboring nodes into the node encoding of the graph.

[0151] Clause 19: At least one non-transient storage medium, wherein, in any one of Clauses 15 to 18, the policy for applying to the outgoing edge is a discrete probability distribution for the outgoing edge at a node of the graph.

[0152] Clause 20: At least one non-transient storage medium, wherein in any one of Clauses 15 to 19, the policy is predicted by training a multilayer perceptron (MLP) using behavior replication.

[0153] In the foregoing description, embodiments of the present invention have been described with reference to numerous specific details that may vary by implementation. Accordingly, the detailed description and drawings should be regarded as illustrative rather than restrictive. The sole proprietary indicator of the scope of the present invention, and what the applicant intends to define as the scope of the present invention, is the literal equivalent of the series of claims appearing in a particular form in this application, including any subsequent amendments. Any definitions expressly described herein for terms included in such claims determine the meaning of such terms used in the claims. Additionally, when the term “further comprising” is used in the foregoing description and in the claims below, what follows this phrase may be an additional step or entity, or a sub-step / sub-entity of the previously mentioned step or entity.

Claims

Claim 1 A method comprising: generating a graph corresponding to a map of a scene by encoding map features and agent features as node encodings of the graph using at least one processor; determining a policy to apply to outgoing edges from nodes of the graph using at least one processor; sampling a path for a target vehicle in the scene according to the policy using at least one processor; predicting a set of trajectories based on the sampled path and sampled latent variables traversed by the policy using at least one processor; and operating a vehicle based on the set of trajectories of the target vehicle using at least one processor. Claim 2 A method according to claim 1, wherein each node corresponds to a segment of the lane centerline of the map. Claim 3 A method according to claim 1 or claim 2, further comprising the step of updating the node encoding to a surrounding agent encoding by calculating a scaled dot product attention weight. Claim 4 A method according to any one of claims 1 to 3, comprising the step of aggregating local context from neighboring nodes into the node encoding of the graph using a graph neural network. Claim 5 A method according to any one of claims 1 to 4, wherein the policy for applying to the outgoing edge is a discrete probability distribution for the outgoing edge at a node of the graph. Claim 6 A method according to any one of claims 1 to 5, wherein the policy is predicted by training a multilayer perceptron (MLP) using behavior replication. Claim 7 A method according to any one of claims 1 to 6, comprising the step of selectively aggregating context along the sampled path and predicting the set of trajectories based on the sampled path traversed by the policy, the aggregated context, and the sampled latent variable. Claim 8 In claim 7, the step of predicting the set of trajectories is: A step of concatenating the aggregated context and the sampled latent variable with motion encoding; and A method comprising the step of inputting the connected aggregated context and the sampled latent variable into a multilayer perceptron, wherein the set of trajectories indicates a predicted position at a future time step. Claim 9 A system comprising: a graph encoder for encoding high-resolution maps and agent features into a graph for generating a final node encoding, wherein the graph includes nodes and edges, wherein the nodes represent segments of lane centerlines and the edges represent transitions between nodes, and the graph is used to generate the final node encoding; a policy header for learning a policy for sampled graph traversal based on the motion of a target vehicle as well as agent context and local scenes at neighboring nodes; and a trajectory decoder for predicting a trajectory based on the node encoding and sampled latent variables along a path traversed by the policy. Claim 10 A system according to claim 9, wherein the policy is a discrete probability distribution of transitions associated with each edge at each node. Claim 11 A system according to claim 9 or claim 10, wherein the graph encoder comprises one or more gated recurrent units for encoding a target vehicle trajectory, surrounding vehicle trajectories, and node features. Claim 12 A system according to any one of claims 9 to 11, wherein the trajectory decoder includes a multihead attention layer that outputs a context vector for each policy, and the context vector is combined with motion encoding and the sampled latent variable to predict the trajectory. Claim 13 A system according to any one of claims 9 to 12, wherein the initial node encoding is updated to a surrounding agent encoding by calculating scaled inner product attention weights to generate the final node encoding. Claim 14 A system according to any one of claims 9 to 13, wherein the graph encoder is configured to use a graph neural network to aggregate local context from neighboring nodes into the final node encoding of the graph. Claim 15 At least one non-transient storage medium storing instructions, wherein when the instructions are executed by at least one processor, the at least one processor: generates a graph corresponding to a map of a scene by encoding map features and agent features as node encodings of the graph; determines a policy to apply to edges outgoing from nodes of the graph; samples a path for a target vehicle in the scene according to the policy; predicts a set of trajectories based on the sampled path and sampled latent variables traversed by the policy; and operates the vehicle based on the set of trajectories of the target vehicle. Claim 16 At least one non-transient storage medium according to claim 15, wherein each node corresponds to a segment of the lane centerline of the map. Claim 17 At least one non-transient storage medium according to claim 15 or claim 16, comprising updating the node encoding to a surrounding agent encoding by calculating a scaled inner attention weight. Claim 18 At least one non-transient storage medium according to any one of claims 15 to 17, comprising using a graph neural network to aggregate local context from neighboring nodes into the node encoding of the graph. Claim 19 At least one non-transient storage medium, wherein, in any one of claims 15 to 18, the policy for applying to the outgoing edge is a discrete probability distribution for the outgoing edge at a node of the graph. Claim 20 At least one non-transient storage medium according to any one of claims 15 to 19, wherein the policy is predicted by training a multilayer perceptron (MLP) using behavior replication.